Configuring the Spectra Detect AMI Scanner
Two settings decide how much a scan uploads and how long it takes: the scan mode, which chooses where the scanner looks, and the file filters, which choose what it sends. This page covers both, along with the Spectra Detect endpoint settings and the tags the scanner applies to the resources it creates.
In a Terraform deployment, each setting below is a variable. A mode or category name the scanner doesn't recognize stops the plan, rather than reaching an instance and being ignored there.
Scan Modes
The scan mode chooses which directories are walked. It's a path filter and nothing more - it never inspects file contents. binary-focused doesn't mean "files that are binaries": a shell script under /usr/bin/ is selected by it, and a compiled executable under /srv/ isn't.
| Mode | Selects |
|---|---|
binary-focused | Executable and library paths only: /bin/, /sbin/, /usr/bin/, /usr/sbin/, /usr/lib/, /usr/lib64/, /lib/, /lib64/, /opt/ |
critical-paths | The executable paths above, plus /etc/, /var/log/, and the SSH directories under /root/ and /home/ |
full-filesystem | Every file the filters don't exclude. This is the default. |
Only these three values are accepted. Anything else is rejected at startup rather than replaced by a default, so a scan never runs under a mode other than the one its report names.
Virtual filesystems - /proc, /sys, /dev, and /run - are excluded before the mode is consulted. No mode walks them.
Set the mode with the scan_mode variable.
File Filters
Filters run after the scan mode, against the files it selected, and they're independent of it. A file is uploaded only if it passes both stages. The mode narrows where the scanner looks; the filters narrow which kinds of file are uploaded. Combining binary-focused with an image exclusion means "executable paths, minus images", not one or the other.
The scanner identifies files by magic bytes, not by extension, so a .txt file holding a JPEG is treated as an image. Detection is skipped entirely when no filter needs it.
Four filters apply, in this order. A file that fails any of them is skipped: it is not uploaded, not analyzed, and is counted in the report's files_skipped. Exclusions are evaluated before allowlists, so a file matching both is skipped:
| Filter | Effect |
|---|---|
scan_exclude_mime | Skip files whose MIME type contains any of these substrings. |
scan_exclude_categories | Skip files in any of these categories. |
scan_include_mime | If set, upload only files whose MIME type contains one of these substrings, skipping the rest. |
scan_include_categories | If set, upload only files in one of these categories, skipping the rest. |
The MIME substring tests aren't anchored. image/ selects image/png because it's a substring of it, and would equally select a type that merely contains image/ later in the string.
Setting both allowlists narrows twice. scan_include_mime and scan_include_categories are evaluated one after the other, so with both set a file must match both to be uploaded - the intersection, not the union. Setting scan_include_categories to document and scan_include_mime to application/zip uploads only files that are both, which for most filesystems is nothing at all.
Categories
Six category names are accepted. An unrecognized name is rejected at startup rather than ignored, because a typo in an exclusion would leave the intended files scanned, and a typo in an allowlist would match nothing at all.
| Category | Covers |
|---|---|
image | PNG, JPEG, GIF, WebP, BMP, TIFF, ICO, SVG |
audio | MP3, FLAC, WAV, OGG, AAC |
video | MP4, MKV, AVI, WebM, MOV |
font | TTF, TTC, OTF |
archive | ZIP, TAR, GZ, BZ2, ZST, 7Z, RAR, XZ |
document | PDF, and legacy OLE2 Office files - DOC, XLS, PPT |
Two things to know about category filtering:
- Modern Office files count as
archive, notdocument. DOCX, XLSX, and PPTX files are ZIP containers, and are indistinguishable from a plain ZIP by magic bytes alone. Excludingarchivetherefore drops them too. - Filtering fails open. When the MIME type can't be determined, the file is uploaded rather than silently skipped. An exclusion reduces scan volume; it isn't a guarantee that no such file is ever sent.
All Filter Settings
| Variable | Sets |
|---|---|
scan_mode | Which directories are walked |
scan_exclude_categories | Categories to skip |
scan_include_categories | Categories to upload, skipping all others |
scan_exclude_mime | MIME substrings to skip |
scan_include_mime | MIME substrings to upload, skipping all others |
scan_exclude_paths | Path patterns to skip |
scan_include_paths | Paths to walk, narrowing the mode further |
scan_file_extensions | Extensions to keep, without the leading period, matched case-insensitively on the filename rather than the contents |
scan_max_files | Cap on files uploaded per scan. 0 means no cap. |
scan_max_file_size | Largest file uploaded, in bytes. 0 means no limit. |
scan_min_file_size | Smallest file uploaded, in bytes |
The defaults leave every filter off, so a deployment applied without any of them scans full-filesystem with no MIME filtering at all - and because no MIME filter is in force, no magic-byte detection runs either.
Filter Examples
Skip media and fonts, which are large relative to their analysis value, and scan everything else:
scan_exclude_categories = ["image", "audio", "video", "font"]
Scan executables and libraries only, and drop archives so a large vendored bundle doesn't dominate the upload:
scan_mode = "binary-focused"
scan_exclude_categories = ["archive"]
Look at documents and archives only, anywhere on the filesystem - an allowlist, so nothing else is uploaded:
scan_include_categories = ["document", "archive"]
Exclude one specific type rather than a whole category, by MIME substring:
scan_mode = "critical-paths"
scan_exclude_mime = ["image/svg"]
Cap what a single scan uploads, which is useful for a first run against an unfamiliar asset:
scan_max_files = 5000
scan_max_file_size = 104857600 # 100 MB
The two caps report differently. A scan that hits scan_max_files is marked partial, carrying the anomaly max_files_truncated: it covers the files the walk reached first, in directory order, rather than the whole asset. Files skipped for exceeding scan_max_file_size are counted like any other filtered file and don't affect the status, so a scan can read complete having never uploaded a large file. See Scan Outcomes and Exit Codes.
Applying a Filter Change
Filters reach an instance through the configuration written at boot, so a change takes effect on instances launched afterward. Applying the change creates a new launch template version and nothing else: the running fleet keeps scanning with its previous filters until those instances are replaced.
Start an instance refresh to apply a change to the fleet immediately. While a refresh is in progress, both filter sets are in force, and a single scan window can produce reports filtered two different ways.
Spectra Detect Endpoint
reversinglabs_api_url is required. It may name a single Spectra Detect Worker, which analyzes files, or a Hub, which ingests them and distributes them to Workers. For the difference, see the Spectra Detect deployment documentation.
Worker Deployments
Pointing reversinglabs_api_url at a Worker directly needs nothing else. The scanner polls only that host.
Hub Deployments
A Hub doesn't analyze files itself. It forwards each submission to one of its connected Workers by round-robin and returns a task URL naming that Worker, and analysis reports can't be obtained from the Hub. The scanner has to poll the Worker directly, so with several Workers behind one Hub, a single scan polls several hosts at once.
Because that host comes from the Hub's response, and because polling it sends the API token, the scanner only follows hosts you've accepted. List them in reversinglabs_allowed_task_hosts:
reversinglabs_api_url = "https://hub.example.com"
reversinglabs_allowed_task_hosts = [".workers.example.com"]
| Entry form | Matches |
|---|---|
worker01.example.com | That host, on any port |
.workers.example.com | Any host in that domain, at any depth - a pool that scales needs no redeployment |
* | Any host the Hub names. Logged as a warning, because it disables the check. |
Matching is by hostname and ignores the port, because the host is what receives the token. A Worker on port 8080 and a Hub on port 443 of the same machine are the same principal. An entry may carry a port for documentation, but that doesn't narrow the match.
A host that isn't listed fails the scan for that file with an explicit error, rather than being polled. If the endpoint URL uses HTTPS, a task URL that uses plain HTTP is refused rather than sending the token in clear text. That applies to every entry form, including *.
Prefer a domain suffix for an autoscaling pool, so Workers added later are accepted without a configuration change. Use * only when the Hub addresses its Workers in a way no suffix can cover - by IP address, for instance - and you're content to treat the Hub as authoritative.
Token Handling
The token reaches the scanner through Secrets Manager, and is never stored as a Terraform variable in state. Either set reversinglabs_api_token and let Terraform create the secret, or set reversinglabs_api_token_secret_arn to reuse a secret you own. Setting both, or neither, fails at plan time.
Resource Tags
Every resource the deployment creates carries a common tag set. Four keys are set for you:
| Tag | Value |
|---|---|
Name | The deployment's name prefix, overridden per resource |
Project | tag_project, which defaults to detect-ami-scanner |
Method | ec2 |
ManagedBy | terraform |
Everything else comes from resource_tags, a free-form map that's empty by default:
resource_tags = {
Environment = "prod"
Owner = "security"
}
Set whatever your account's tag policy and cost allocation scheme require.
Project and ManagedBy aren't safe to redefine. Project is what grants the scanner permission to attach and detach volumes on its own instances, and ManagedBy is what distinguishes a resource created by Terraform from one created by the scanner at run time.
An Environment key inside resource_tags is unrelated to the environment variable. The latter feeds the name prefix, and so appears in every resource name.
Tags on Resources the Scanner Creates
The snapshots and volumes the scanner creates during a scan aren't Terraform resources, so they can't inherit the common tag set directly. The same set is passed to the scanner instead, and the scanner merges your tags underneath its own.
The scanner's own tags are always applied and can't be overridden:
| Tag | Value |
|---|---|
scanner:managed | true |
ManagedBy | ami-scanner |
CreatedAt | Creation time |
SourceSnapshot, SourceVolume, or SnapshotID | The source the resource was created from |
A configured tag loses a key collision. A configured ManagedBy doesn't displace ami-scanner on a scanner-created resource, which is what keeps a single tag query returning everything the scanner made. A configured scanner:managed is ignored rather than rejected, so you can pass a deployment-wide tag set through without filtering keys out of it.
Tags are validated when the configuration loads: at most 46 configured tags, because EC2 caps a resource at 50 and the scanner adds up to four of its own; keys up to 128 characters; values up to 256; and no reserved aws: prefix. A malformed pair or a limit breach stops the scanner at startup and names the offender, rather than surfacing later as a denied resource creation that reads like a missing permission.
Tags aren't copied from the asset under scan. A temporary volume is the scanner's own resource, not a derivative of the target, and inheriting a production image's owner or cost center would misattribute it. An Amazon-owned or cross-account source has no tags to inherit in any case.